Proxies & Business
October 9, 2026
7 min

Headless browser in Python: A practical guide (Selenium, Playwright, Pyppeteer + proxies)

Alex Sadovskij
Alex Sadovskij
CEO Proxy-Cheap
Headless browser in Python: A practical guide (Selenium, Playwright, Pyppeteer + proxies)
Краткое содержание
Run a headless browser in Python with Selenium, Playwright, or Pyppeteer. Tested code, authenticated proxy setup for each library, and a speed comparison.

Key takeaways

  • A headless browser runs a real browser engine without a window, so your script gets the rendered HTML that plain HTTP requests miss on JavaScript-heavy pages.
  • Playwright is the best default for new projects. In our 20-URL test, it took 7.4 seconds versus 10.2 for Selenium, and 5.4 with async pages.
  • Every library takes a proxy at launch, but logins differ. Selenium 4 uses a WebDriver BiDi auth handler, Playwright a proxy dict, and Pyppeteer page.authenticate().
  • Match the proxy to the job: rotating residential for high-volume collection, static residential (ISP) for session-bound work, datacenter IPv4 for fast runs on public pages.

What is a headless browser in Python?

A headless browser is a real browser engine, such as Chrome, Firefox, or WebKit, that runs without a visible window. In Python, libraries like Selenium, Playwright, and Pyppeteer let you load pages, run JavaScript, and access the fully rendered HTML that a basic HTTP request may miss.

This matters for single-page apps where content only appears after JavaScript executes. A headless browser builds the same DOM as a regular browser, but without displaying it, making it ideal for servers, CI pipelines, and scheduled jobs.

Python’s main options are Selenium, Playwright, and Pyppeteer. All three support proxy connections, which we’ll cover later.

How to run a headless browser in Python with Selenium

To run headless Chrome with Selenium in Python, create a ChromeOptions object, add the --headless=new argument, and pass it to webdriver.Chrome. Selenium Manager, built into Selenium 4.6 and later, fetches the matching driver automatically. Call driver.get(url), wait for an element with WebDriverWait, then read driver.page_source or extract elements by CSS selector.

We tested everything below with Selenium 4.49.0, Python 3.11, and Chrome 148.

bash pip install selenium

Older tutorials add the webdriver-manager package, but you don't need it anymore. Selenium Manager finds your Chrome version and downloads the right ChromeDriver.

This script opens the JavaScript version of the Quotes to Scrape sandbox, waits for the quotes to render, and prints them:

python from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC options = webdriver.ChromeOptions() options.add_argument("--headless=new") options.add_argument("--window-size=1920,1080") driver = webdriver.Chrome(options=options)  # Selenium Manager resolves the driver try:    driver.get("https://quotes.toscrape.com/js/")    WebDriverWait(driver, 10).until(        EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.quote"))    )    for quote in driver.find_elements(By.CSS_SELECTOR, "div.quote"):        text = quote.find_element(By.CSS_SELECTOR, "span.text").text        author = quote.find_element(By.CSS_SELECTOR, "small.author").text        print(f"{author}: {text}") finally:    driver.quit()

A few details matter here:

  • --headless=new vs --headless. Chrome 132 removed the old headless mode from the main binary, so both flags now start the new mode. Keeping =new is explicit and still works on older installs.
  • Wait, don't sleep. WebDriverWait polls until the elements exist. A fixed time.sleep(5) either wastes time or fails on a slow page.
  • Always quit. The try/finally pattern closes Chrome even when something fails, so no orphan processes eat RAM on your server.

To parse with BeautifulSoup instead, pass it driver.page_source. For the networking side, our explainer on the proxy API covers how proxies expose endpoints to code. The full reference is in the Selenium documentation.

How to run a headless browser in Python with Playwright

Install Playwright with pip install playwright, then playwright install chromium. Launch with p.chromium.launch(headless=True), open a page, and call page.goto(url). Playwright waits for elements automatically, so you can call page.locator(...) without manual sleeps. It drives Chromium over the Chrome DevTools Protocol, which makes it faster than Selenium's WebDriver path.

bash pip install playwright playwright install chromium # On a fresh Linux server, add system libraries too: # playwright install --with-deps chromium

Playwright downloads its own browser builds, so there's no driver to match. Current releases need Python 3.10 or later, and we tested with Playwright 1.63.0.

Here's the same quotes job with the sync API:

python from playwright.sync_api import sync_playwright with sync_playwright() as p:    browser = p.chromium.launch(headless=True)    page = browser.new_page()    page.goto("https://quotes.toscrape.com/js/")    quotes = page.locator("div.quote")    quotes.first.wait_for()  # explicit wait for JS-rendered content    for i in range(quotes.count()):        text = quotes.nth(i).locator("span.text").inner_text()        author = quotes.nth(i).locator("small.author").inner_text()        print(f"{author}: {text}")    page.screenshot(path="quotes.png", full_page=True)    browser.close()

Locator actions like inner_text() and click() wait for the element on their own. count() doesn't, which is why the script waits on quotes.first before counting. The screenshot line is a one-liner, and page.pdf(path="quotes.pdf") does the same for PDFs in Chromium.

The async API is where Playwright pulls ahead, because one browser can load several pages at once:

python import asyncio from playwright.async_api import async_playwright URLS = [f"https://quotes.toscrape.com/js/page/{n}/" for n in range(1, 6)] async def scrape(browser, url):    page = await browser.new_page()    await page.goto(url)    await page.locator("div.quote").first.wait_for()    count = await page.locator("div.quote").count()    await page.close()    return url, count async def main():    async with async_playwright() as p:        browser = await p.chromium.launch(headless=True)        results = await asyncio.gather(*(scrape(browser, u) for u in URLS))        for url, count in results:            print(url, count)        await browser.close() asyncio.run(main())

By default, headless=True uses a lightweight Chromium headless shell. Pass channel="chromium" to launch() if you want a full, new headless Chrome instead. Pick Playwright for new projects, async jobs, or when you need Chromium, Firefox, and WebKit from one API. It also pairs well with rotating residential proxies for multi-page jobs. More options are in the Playwright Python documentation.

How to run a headless browser in Python with Pyppeteer

Pyppeteer is a Python port of Puppeteer that drives Chromium through the DevTools Protocol and is async by default. In 2026, use pyppeteer==2.0.0 in its own virtualenv. It requires websockets 10.x and can conflict with Selenium 4’s urllib3 2.x dependency.

Older tutorials often pin websockets==8.1 for Pyppeteer 0.x, so follow current version guidance. Our install tests found:

  • Clean virtualenv: pyppeteer 2.0.0 installed websockets 10.4 and urllib3 1.26.20 and ran correctly.
  • With Selenium 4.49: pip downgraded to pyppeteer 0.0.25, which crashed with AttributeError: module 'websockets' has no attribute 'client'.

Use a separate environment and pin the version:

bash python -m venv pyppeteer-env source pyppeteer-env/bin/activate pip install pyppeteer==2.0.0python import asyncio from pyppeteer import launch async def main():    browser = await launch(headless=True, args=["--no-sandbox"])    page = await browser.newPage()    await page.goto("https://quotes.toscrape.com/js/")    await page.waitForSelector("div.quote")    quotes = await page.querySelectorAllEval(        "div.quote",        "nodes => nodes.map(n => n.querySelector('small.author').innerText"        " + ': ' + n.querySelector('span.text').innerText)",    )    for line in quotes:        print(line)    await browser.close() asyncio.run(main())

On first launch, Pyppeteer downloads a Chromium build of about 100MB, or you can use an existing browser with executablePath. Recent Chrome versions may print a harmless “Future exception was never retrieved” message on close.

Pyppeteer is Chromium-only and no longer actively maintained. Its README recommends Playwright, and the latest release is 2.0.0 from February 2024. It still works well for porting Puppeteer scripts, but the Pyppeteer documentation shows the older 0.0.25 version, so check PyPI for current pins.

How to route a headless browser through a proxy in Python

Each library accepts a proxy at launch. Selenium uses --proxy-server=host:port with ChromeOptions and a BiDi auth handler. Playwright accepts proxy={"server", "username", "password"} at launch. Pyppeteer uses --proxy-server and page.authenticate.

Residential or ISP proxies in your target market help avoid per-IP rate limits and ensure region-specific pages match what local visitors see. This is useful for localization testing, price checks, and QA on public content.

The snippets use placeholder credentials. Get your real values from the Proxy-Cheap dashboard, where you can select a country and choose rotating IPs or sticky sessions of about 30 minutes. We tested all three over HTTPS through a local password-protected HTTP proxy.

Playwright is the simplest. The proxy is a dict passed to launch():

python from playwright.sync_api import sync_playwright PROXY = {    "server": "http://your-proxy-host:your-proxy-port",    "username": "your-username",    "password": "your-password", } with sync_playwright() as p:    browser = p.chromium.launch(headless=True, proxy=PROXY)    page = browser.new_page()    page.goto("https://httpbin.org/ip")    print(page.inner_text("body"))    browser.close()

Selenium takes the proxy address as a Chrome flag, but Chrome won't accept a username and password there. Selenium 4 handles the login through WebDriver BiDi:

python from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC PROXY_HOST = "your-proxy-host"   # copy these from the Proxy-Cheap dashboard PROXY_PORT = "your-proxy-port" PROXY_USER = "your-username" PROXY_PASS = "your-password" options = webdriver.ChromeOptions() options.add_argument("--headless=new") options.add_argument(f"--proxy-server=http://{PROXY_HOST}:{PROXY_PORT}") options.enable_bidi = True  # WebDriver BiDi handles the proxy login driver = webdriver.Chrome(options=options) driver.network.add_auth_handler(PROXY_USER, PROXY_PASS) def open_page(url):    # Navigate over BiDi so the auth handler can answer the proxy challenge    driver.browsing_context.navigate(        context=driver.current_window_handle, url=url, wait="complete"    ) try:    open_page("https://httpbin.org/ip")    print(driver.find_element(By.TAG_NAME, "body").text)  # exit IP    open_page("https://quotes.toscrape.com/js/")    WebDriverWait(driver, 10).until(        EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.quote"))    )    print(len(driver.find_elements(By.CSS_SELECTOR, "div.quote")), "quotes") finally:    driver.quit()

One gotcha from our testing: with the auth handler active, a plain driver.get() stalled until the page load timeout on Selenium 4.49 and Chrome 148. Classic navigation waits for the page, while the proxy login waits on a BiDi reply. Navigating with driver.browsing_context.navigate() keeps both on BiDi, and the page loaded in under a second. On IP whitelist plans, skip the handler and keep only the --proxy-server flag.

Pyppeteer takes the flag in args and the login through page.authenticate():

python import asyncio from pyppeteer import launch PROXY_HOST = "your-proxy-host" PROXY_PORT = "your-proxy-port" PROXY_USER = "your-username" PROXY_PASS = "your-password" async def main():    browser = await launch(        headless=True,        args=["--no-sandbox", f"--proxy-server=http://{PROXY_HOST}:{PROXY_PORT}"],    )    page = await browser.newPage()    await page.authenticate({"username": PROXY_USER, "password": PROXY_PASS})    await page.goto("https://httpbin.org/ip")    print(await page.evaluate("document.body.innerText"))    await browser.close() asyncio.run(main())

Rotating residential proxies use username and password, built for high-volume rotation. To keep credentials out of code, static residential (ISP) proxies and datacenter plans also support IP whitelist authentication. See our web data collection page for related workloads, and the Proxy-Cheap API docs for managing proxies from code.

Which Python proxy pairs with which headless workload

Match the proxy to the job. Use rotating residential for high-volume, distributed collection where each request can come from a fresh IP. Use static residential (ISP) for account-bound sessions that must keep one identity. Use datacenter IPv4 for high-throughput work on publicly available, unprotected pages where speed and cost matter most.

WorkloadRecommended Proxy-Cheap productAuthWhy it fits
High-volume collection across many pagesRotating residentialUsername/passwordNew IP per request or a sticky session of about 30 minutes, 180+ locations, pay-as-you-go per GB
Account-bound or logged-in sessionsStatic residential (ISP)Username/password or IP whitelistOne fixed IP for the whole subscription, country and ISP targeting, unlimited bandwidth
High-throughput runs on public pagesDatacenter IPv4 (IPv6 also available)Username/password or IP whitelistPer-IP pricing, unlimited bandwidth, fixed IPs for repeatable daily crawls
Mobile-first sites and mobile QARotating mobileIP whitelistCarrier IPs from 5G, 4G, and LTE networks in 100+ countries

Rotating residential has no limit on concurrent sessions, which suits async Playwright jobs. Static residential and datacenter plans are sized for about 100 concurrent connections per proxy. With datacenter proxies, per-IP pricing keeps costs predictable for a crawl you run every day.

4G and 5G mobile proxies show you what visitors on mobile networks see, which pairs well with Playwright's device emulation. For background on each type, read our guide to residential vs datacenter proxies.

Ready to test your setup? Proxy-Cheap runs on pay-as-you-go billing with no monthly commitment, so you can try a small plan against your own script before scaling. For session-heavy headless work, start with static and rotating ISP proxies and drop the credentials into the snippets above.

Selenium vs Playwright vs Pyppeteer: a quick comparison

Playwright is the fastest and most modern, with automatic waiting and native async support. Selenium supports the widest browser range and has the largest community. Pyppeteer is a lightweight, Chromium-only choice for porting Puppeteer code. For most new Python projects, start with Playwright; use Selenium when browser coverage matters.

LibraryBrowsersSpeed (20 URLs)AsyncBest for
Playwright 1.63Chromium, Firefox, WebKit7.4s sync, 5.4s asyncYesNew projects
Selenium 4.49Chrome, Firefox, Edge, Safari10.2sNo native asyncCross-browser testing
Pyppeteer 2.0.0ChromiumNot benchmarkedAsync onlyPorting Puppeteer

Benchmark: Selenium and Playwright loaded 20 JavaScript-rendered pages and extracted 10 quotes from each on a 2-vCPU, 8GB Linux sandbox. Figures are the median of three runs, including browser startup. Playwright was about 1.4x faster sequentially and 1.9x faster with five concurrent pages.

Real sites will take longer, but the ranking should hold. Speed matters most for recurring jobs such as SEO research workflows.

Memory matters too. One headless Chrome instance used about 310MB in our sandbox, within the reported 200 to 500MB range.

Can a website detect a headless browser?

Yes. Sites can detect automation through signals like navigator.webdriver, missing browser features, unusual request patterns, and the "HeadlessChrome" user agent. Hundreds of requests per minute from one IP can also stand out.

For reliable, location-accurate sessions, use residential or ISP IPs that match your target market and collect only public data. Be a considerate client:

  • Pace requests: Add delays and limit concurrency per domain.
  • Respect robots.txt and rate limits: Check both before crawling.
  • Back off on 429/503: Honor Retry-After and slow down when errors increase.
  • Match the market: Use residential or ISP IPs in the country you're testing.
  • Use public data only: Avoid unauthorized logins and collecting personal data without a valid basis.

What are the disadvantages of a headless browser?

Headless browsers are heavier than plain HTTP requests. Each instance uses roughly 200-500MB of RAM and adds time to every page, so a single server handles only a handful of concurrent instances. They also need browser dependencies that are easy to miss on a fresh server, which is why scripts that work locally fail there.

The same Firecrawl analysis estimates a 4GB server handles about 5 to 8 concurrent browser instances. It puts a self-hosted headless setup at $190 to $780 a month, covering compute, proxies, storage, and monitoring.

Server fragility is the second pain point. These fixes come up most often:

  • Missing system libraries. Run playwright install --with-deps chromium on Linux, or install Chrome's dependencies with your package manager.
  • Docker quirks. Chrome often needs --no-sandbox when running as root in a container, plus --disable-dev-shm-usage if /dev/shm is small.
  • Driver mismatches. ChromeDriver's major version must match Chrome's. Selenium Manager handles this unless an old pinned driver sits on the server.

Sometimes the best fix is no browser at all. If the data is in the initial HTML, or loads from a JSON endpoint you can call directly, requests plus BeautifulSoup is faster and cheaper. Check the page source and network tab first.

Is Selenium still relevant in Python?

Yes. Selenium remains widely used in 2026 because it supports Chrome, Firefox, Safari, and Edge, has the largest community, and integrates with most test frameworks. Playwright is faster for new scraping projects, but Selenium is still the pragmatic pick when you need broad browser coverage or existing test infrastructure.

It keeps improving, too. Selenium Manager removed most driver headaches, and WebDriver BiDi now handles proxy logins that used to need extra packages.

Часто задаваемые вопросы

For most new projects, Playwright. It has auto-waiting, native async, and bundled browsers, and it was the fastest in our test. Choose Selenium for Safari or Edge coverage or an existing test suite, and Pyppeteer for porting Puppeteer code.

Not for small local tests. For volume or location-accurate work, yes. One server IP runs into per-IP rate limits and only sees one market's version of a page. Residential or ISP IPs in your target country give you stable sessions and matching data.

Headful mode opens a visible browser window. Headless mode runs the same engine and builds the same page without drawing anything. Headless suits servers and CI, while headful helps with debugging because you can watch each step.

Yes, when the content is in the initial HTML or comes from a JSON endpoint you can call directly. In those cases, requests plus BeautifulSoup is faster and lighter.

Usually it's missing system libraries, a Docker sandbox issue, or a browser and driver mismatch. The server's single IP can also hit rate limits or return region-specific content, which a proxy in your target market solves.

Playwright, in our test. It finished 20 pages in 7.4 seconds (5.4 with async) versus 10.2 seconds for Selenium.

No. Playwright downloads its own browser builds when you run playwright install. Selenium still uses a driver such as ChromeDriver, but Selenium Manager now downloads it for you.

Pass a proxy dict with server, username, and password keys to p.chromium.launch(). You can also set one per context with browser.new_context(proxy=...) when different pages need different IPs.

Its README says it's no longer actively maintained and recommends Playwright, and the latest release is 2.0.0 from February 2024. It still works for porting Puppeteer code if you install it in its own virtualenv.